Papers by Sabrina J. Mielke

11 papers
Reducing Conversational Agents’ Overconfidence Through Linguistic Calibration (2022.tacl-1)

Copied to clipboard

Challenge: Neural generative open-domain english-language dialogue agents are currently unsuitable for applications other than entertainement and research.
Approach: They propose to incorporate metacognitive features into the training of a controllable generation model to improve likelihood of correctness.
Outcome: The proposed model improves likelihood of correctness by incorporating metacognitive features into the training of a controllable generation model.
Unsupervised Disambiguation of Syncretism in Inflected Lexicons (N18-2)

Copied to clipboard

Challenge: Lexical ambiguity makes it difficult to compute useful statistics of a corpus.
Approach: They propose a neural network-based model that fits a prior distribution over feature bundles to a list of unigram type counts and partitions each count among different analyses of that unigrammer.
Outcome: The proposed model is based on a list of unigram type counts and partitions each count among different analyses of that unigrammer.
What Kind of Language Is Hard to Language-Model? (P19-1)

Copied to clipboard

Challenge: a recent study suggests that language models perform poorly across languages.
Approach: They propose a model that fits a paired-sample multiplicative mixed-effects model to obtain language difficulty coefficients from at least-pairwise parallel corpora.
Outcome: The proposed model is able to handle missing data and is aware of inter-sentence variation.
Counterfactual Data Augmentation for Mitigating Gender Stereotypes in Languages with Rich Morphology (P19-1)

Copied to clipboard

Challenge: Gender stereotypes are manifest in most of the world's languages and are consequently propagated or amplified by NLP systems.
Approach: They propose a method for converting between masculine-inflected and feminine-infflectes sentences in morphologically rich languages to reduce gender stereotyping by a factor of 2.5 without any sacrifice to grammaticality.
Outcome: The proposed approach reduces gender stereotyping by 2.5 without any sacrifice to grammaticality.
A Structured Variational Autoencoder for Contextual Morphological Inflection (P18-1)

Copied to clipboard

Challenge: morphological inflectors typically trained on fully supervised, type-level data, but how can we improve their performance? et al., 2016: a novel latent-variable model for semi-supervised learning of inflection generation.
Approach: They propose a latent-variable model for semi-supervised learning of inflection generation . they use a wake-sleep algorithm to enable posterior inference over latent variables .
Outcome: The proposed model improves on 23 languages and shows 10% accuracy improvement . the proposed model is based on the wake-sleep algorithm .
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset (2020.lrec-1)

Copied to clipboard

Challenge: a new resource is available for 12 South Asian languages that use the Latin script for text entry . the Latin-script system is not widely used in South Asian language writing, despite the Latin alphabet .
Approach: They describe the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages.
Outcome: The Dakshina dataset includes text in both the Latin and native scripts for 12 languages . the authors provide baseline results on several tasks made possible by the dataset .
UniMorph 2.0: Universal Morphology (L18-1)

Copied to clipboard

Challenge: The Universal Morphology project is a collaborative effort to improve how NLP handles complex morphology across the world's languages.
Approach: They propose to use a universal tagset to annotate morphological data using a schema that includes a lemma and a bundle of morphology features.
Outcome: The project releases annotated morphological data using a universal tagset, the UniMorph schema.
Are All Languages Equally Hard to Language-Model? (N18-2)

Copied to clipboard

Challenge: a fair comparison of language models is tricky because of the size of the corpora and the variability of orthographic systems.
Approach: They propose a framework for fair cross-linguistic comparison of language models . they show that in some languages, textual expression is harder to predict with n-gram models compared to LSTM models based on translated text .
Outcome: The proposed framework is based on translated text and language models on 21 languages.
Tired of Topic Models? Clusters of Pretrained Word Embeddings Make for Fast and Good Topics too! (2020.emnlp-main)

Copied to clipboard

Challenge: Existing topic models rely on probabilistic models to uncover themes within document collections, but are they the only option?
Approach: They propose a way to cluster pre-trained word embeddings while incorporating document information for weighted clustering and reranking top words.
Outcome: The proposed approach performs as well as classical topic models, but with lower runtime and computational complexity.
UniMorph 3.0: Universal Morphology (2020.lrec-1)

Copied to clipboard

Challenge: Explicit modeling of morphology has demonstrable benefits for language modeling, speech recognition, word embedding and keyword search.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource for annotated data in diverse languages.
Outcome: The proposed schema has been improved to make it more complete and correct, and adds 66 new languages and parts of speech for 12 languages.
It’s Easier to Translate out of English than into it: Measuring Neural Translation Difficulty by Cross-Mutual Information (2020.acl-main)

Copied to clipboard

Challenge: Current state-of-the-art MT systems are based on neural networks, but it is unclear whether all translation directions are equally easy (or hard) to model for NMT.
Approach: They propose an asymmetric information-theoretic metric of machine translation difficulty that exploits the probabilistic nature of most neural machine translation models.
Outcome: The proposed metric allows us to better evaluate the difficulty of translating text into the target language while controlling for the difficulty independent of the translation task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations